Papers with out-of-vocabulary word embedding
Glyph2Vec: Learning Chinese Out-of-Vocabulary Word Embedding from Glyphs (2020.acl-main)
Copied to clipboard
| Challenge: | Chinese NLP applications that rely on large text often contain huge amounts of vocabulary which are sparse in corpus. |
| Approach: | They propose a multi-modal model that extracts visual features from Chinese word glyphs to expand current word embedding space without accessing any corpus. |
| Outcome: | The proposed model can embed words in Chinese without accessing corpus without a corpus. |